Using Grammatical Inference to Improve Precision in Information Extraction

نویسنده

  • Dayne Freitag
چکیده

The eld of information extraction (IE) is concerned with applying natural language processing (NLP) and information retrieval (IR) techniques to the automatic extraction of essential details from text documents. We are exploring the use of machine learning methods for IE. While the most promising methods we have developed perform well for problems deened over a collection of electronic seminar announcements, they are imprecise in their identiication of the boundaries of relevant text fragments (elds). Here, we entertain the idea of using grammatical inference (GI) methods to learn the appropriate form of a eld. We describe one method for translating raw text into an abstract alphabet suitable for GI, and show that, by combining one IE learning method with the resulting inferred grammars, large improvements in precision can be realized for some elds.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Information Extraction using Context-free Grammatical Inference from Positive Examples

Information extraction from textual data has various applications, such as semantic search. Learning from positive example have theoretical limitations, for many useful applications (including natural languages), substantial part of practical structure (CFG) can be captured by framework introduced in this paper. Our approach to automate identification of structural information is based on gramm...

متن کامل

A New Method for Improving Computational Cost of Open Information Extraction Systems Using Log-Linear Model

Information extraction (IE) is a process of automatically providing a structured representation from an unstructured or semi-structured text. It is a long-standing challenge in natural language processing (NLP) which has been intensified by the increased volume of information and heterogeneity, and non-structured form of it. One of the core information extraction tasks is relation extraction wh...

متن کامل

Extracting Key Terms from Chinese and Japanese texts

Key term extraction is very useful for information retrieval. Most term extraction methods use one of two approaches, namely lexical and grammatical. We argue that due to the diierences in linguistic and character set characteristics of Chinese and Japanese, a lexical approach is more suitable for Chinese whereas a grammatical approach is more suitable for Japanese. In this paper, we present tw...

متن کامل

More Informative Open Information Extraction via Simple Inference

Recent Open Information Extraction (OpenIE) systems utilize grammatical structure to extract facts with very high recall and good precision. In this paper, we point out that a significant fraction of the extracted facts is, however, not informative. For example, for the sentence The ICRW is a non-profit organization headquartered in Washington, the extracted fact (a non-profit organization) (is...

متن کامل

Towards Schema-Guided XML Query Induction

XML query induction is a key task in Web information extraction. Recent approaches based on grammatical inference represent node selection queries in XML trees by deterministic tree automata. In this paper, we show how to guide RPNI-based learning algorithms by XML schemas which we can infer in a preprocessing step. We hope that schema guidance will help to improve heuristics that are essential...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 1997